An kNN Model-Based Approach and Its Application in Text Categorization
نویسندگان
چکیده
An investigation has been conducted on two well known similarity-based learning approaches to text categorization. This includes the k-nearest neighbor (kNN) classifier and the Rocchio classifier. After identifying the weakness and strength of each technique, we propose a new classifier called the kNN model-based classifier by unifying the strengths of k-NN and Rocchio classifier and adapting to characteristics of text categorization problems. A text categorization prototypes system has been implemented and then evaluated on two common document corpora, namely, the 20-newsgroup collection and the ModApte version of the Reuters-21578 collection of news stories. The experimental results show that the kNN model-based approach outperforms the kNN, Rocchio classifier.
منابع مشابه
Multiclass Boosting with Adaptive Group-Based kNN and Its Application in Text Categorization
AdaBoost is an excellent committee-based tool for classification. However, its effectiveness and efficiency in multiclass categorization face the challenges from methods based on support vector machine SVM , neural networks NN , naı̈ve Bayes, and k-nearest neighbor kNN . This paper uses a novel multi-class AdaBoost algorithm to avoid reducing the multi-class classification problem to multiple tw...
متن کاملUsing kNN Model-based Approach for Automatic Text Categorization
An investigation has been conducted on two well known similarity-based learning approaches to text categorization: the k-nearest neighbor (k-NN) classifier and the Rocchio classifier. After identifying the weakness and strength of each technique, a new classifier called the kNN model-based classifier (kNNModel) has been proposed. It combines the strength of both k-NN and Rocchio. A text categor...
متن کاملA Summary Writing Model Based on Van Dijk’s Concept of Macrostructure and its Application within the Genre-Based Approach
This study was an attempt to provide a comprehensive model for summary writing based on the model of Van Dijk’s concept of macrostructures. The effectiveness of the model was examined in a genre-based quasi-experimental study with the data collection procedure lasting a semester. The participants included 60 female English learners divided into two experimental and control groups. The results o...
متن کاملInverted Index based Modified Version of KNN for Text Categorization
This research proposes a new strategy where documents are encoded into string vectors and modified version of KNN to be adaptable to string vectors for text categorization. Traditionally, when KNN are used for pattern classification, raw data should be encoded into numerical vectors. This encoding may be difficult, depending on a given application area of pattern classification. For example, in...
متن کاملA ME Model Based on Feature Template for Chinese Text Categorization
With entering into information society and the Internet developing rapidly, people could acquire more and more information. How to utilize Internet information efficiently and promptly, has became a hotspot in information technology. Text categorization is an important component to help getting useful message from tremendous amount of vast information. And it assigns new documents to pre-define...
متن کامل